Papers with safety metrics

6 papers
Evaluating Psychological Safety of Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: a recent study evaluated the psychological safety of large language models.
Approach: They designed unbiased prompts to evaluate the psychological safety of large language models.
Outcome: The proposed prompts showed that they were fine-tuned with behavioral metrics to reduce toxicity.
VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment (2024.emnlp-main)

Copied to clipboard

Challenge: Large vision-language models (LVLMs) are evolving rapidly and require data with human supervision to achieve better alignment.
Approach: They introduce VLFeedback, the first large-scale vision-language feedback dataset . they train Silkie, an LVLM fine-tuned via direct preference optimization .
Outcome: The proposed model outperforms its base model in helpfulness, visual faithfulness, and safety metrics and exhibits enhanced resilience against red-teaming attacks.
Soteria: Language-Specific Functional Parameter Steering for Multilingual Safety Alignment (2025.findings-emnlp)

Copied to clipboard

Challenge: Soteria locates and minimally adjusts the “functional heads” most responsible for harmful content generation in each language.
Approach: Soteria locates and minimally adjusts the "functional heads" responsible for harmful content generation in each language.
Outcome: The proposed approach reduces harmful content generation in languages while preserving model performance.
The Rise of Darkness: Safety-Utility Trade-Offs in Role-Playing Dialogue Agents (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) demonstrate their utility in character simulations, but they pose a risk of generating unsafe content.
Approach: They propose a method which dynamically adjusts safety-utility preferences based on the degree of risk coupling and guides the model to generate responses biased toward utility or safety.
Outcome: The proposed method improves safety metrics while maintaining utility.
Sowing the Wind, Reaping the Whirlwind: The Impact of Editing Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) face challenges in maintaining accuracy due to the dynamic nature of world knowledge.
Approach: They propose to use a benchmark dataset to investigate the effects of model edits on model safety metrics and guardrails.
Outcome: The proposed dataset sheds light on how the edits, impact the model’s safety metrics and guardrails.
Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study (2025.emnlp-main)

Copied to clipboard

Challenge: Existing safety evaluations rely on artificial images to evaluate vision-language models . a recent study found that memes are more effective at bypassing safety measures than synthetic or typographic images.
Approach: They propose a benchmark pairing meme images with harmful and benign instructions . they assess multiple VLMs across single and multi-turn interactions .
Outcome: The proposed benchmark pairs real meme images with harmful and benign instructions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations